Papers by Marcely Zanon Boito
Speech Resources in the Tamasheq Language (2022.lrec-1)
Copied to clipboard
Marcely Zanon Boito, Fethi Bougares, Florentin Barbier, Souhir Gahbiche, Loïc Barrault, Mickael Rouvier, Yannick Estève
| Challenge: | In this paper, we present two datasets for Tamasheq, a developing language mainly spoken in Mali and Niger . we share unlabeled audio data in five languages: french, Fulfulde, Hausa, Tamaheq and Zarma . |
| Approach: | They present two datasets for Tamasheq, a developing language mainly spoken in Mali and Niger. |
| Outcome: | The proposed datasets are used in the IWSLT 2022 low-resource speech translation track . they consist of radio recordings from daily broadcast news in Niger and Mali . |
From Tower to Spire: Adding the Speech Modality to a Translation-Specialist LLM (2025.findings-emnlp)
Copied to clipboard
Kshitij Ambilduke, Ben Peters, Sonal Sannigrahi, Anil Keshwani, Tsz Kin Lam, Bruno Martins, Andre Martins, Marcely Zanon Boito
| Challenge: | Spire is a speech-augmented language model capable of translating speech input into 10 languages and transcribing text input in both directions. |
| Approach: | They introduce a speech-augmented language model capable of translating speech input into 10 languages . they integrate the model into existing multilingual LMs via speech discretization . |
| Outcome: | Spire integrates speech-augmented language model into existing multilingual model using speech discretization and pre-training using only 42.5 K hours of speech. |
MaSS: A Large and Clean Multilingual Corpus of Sentence-aligned Spoken Utterances Extracted from the Bible (2020.lrec-1)
Copied to clipboard
| Challenge: | The Bible is the same for all the languages, thus constituting a multilingual and comparable 2 spoken corpus, is not exploited to date. |
| Approach: | They propose to add multilingual links between small speech segments in different languages . they use a large dataset of 8,130 parallel spoken utterances across 8 languages - maSS . |
| Outcome: | The proposed model can build automatic speech recognition models for 700 languages. |